Skip to content

Add the whiteboard interview mode - #72

Merged
jserv merged 9 commits into
sysprog21:mainfrom
ColtenOuO:feat/whiteboard-mode
Oct 5, 2026
Merged

jserv merged 9 commits into
sysprog21:mainfrom
ColtenOuO:feat/whiteboard-mode

Conversation

@ColtenOuO

@ColtenOuO ColtenOuO commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A whiteboard interview is the coding round held without a compiler. The lobby offers it next to the duration and the loop. The editor, the starter code and the test runner are replaced by a drawing surface, and the candidate shows correctness by drawing an approach, tracing an example across it, and naming the cases that would break it. It uses the same problem bank, the same six REACTO phase ids, the same rubric and the same report schema. Only the surface and what each step asks for change.

Closes #65, with two deliberate departures from the issue's proposal:

  • No handwritten pseudo-code step. The flow is the six REACTO steps, so the evidence vocabulary, the progress checklist and the rubric all stay as they are.
  • No new grading dimensions. The report is graded against the existing rubric. The reviewer is told what each step means at a board.

What the candidate sees

  • Board. A fixed 1600x1000 canvas that scrolls inside its panel, so the interviewer always sees the whole board whatever the candidate's screen size. Mouse input only in this version.
  • Pens. Four colours, shown on a white strip so each swatch looks like the ink it draws.
  • Eraser. Three sizes: 12, 24 and 56 board px. The cursor becomes a circle the size of the erasure.
  • Undo, redo and Clear board. Clear is a single undoable step.
  • Opening. The greeting does not ask for a language, since there are no tabs and nothing compiles. It goes straight to restating the problem.
  • Checklist. The step list beside the timer uses the whiteboard names below.
Phase id At the board What is asked
repeat Repeat restate inputs, outputs, constraints and ambiguities
example Example draw one ordinary case and one boundary case
algorithm Approach draw the approach and its cost before tracing it
coding Trace walk one of their examples through the drawing
test Edge cases name what would break it, and what it does on each
optimizations Complexity confirm time and space, name one optimization

How a board reaches the interviewer

The browser keeps the drawing as vector strokes (web/whiteboard.js). That is what makes undo and replay possible. Rendering a board to pixels happens only when something needs an image.

flowchart LR
    Stroke["Stroke ends"] --> Settle["1 s with no new stroke"]
    Settle --> JPEG["Canvas exported as JPEG<br/>quality 0.72"]
    JPEG -->|"LiveKit byte stream<br/>topic board_image"| Drain["Agent drain task<br/>bounded at 512 KiB"]
    Drain -->|"queue of 2"| Loop["Room loop: pump_board"]
    Loop --> Latest["Latest board in memory"]
    Loop -->|"at most one per second<br/>realtimeInput.video"| Live["Gemini Live"]
    Tool["read_board tool call"] --> Resend["Resend latest board"]
    Latest --> Resend
    Resend --> Live
Loading
  • Why a byte stream. A board is tens of kilobytes, more than one data packet carries. The stream header carries a strokes count and, for a phase checkpoint, a checkpoint id.
  • What the agent refuses. A stream is refused unless it is on the board topic, from the candidate, in a whiteboard interview, and declares no more than 512 KiB. The same bound is enforced again while reading, for a sender that declared nothing.
  • Keeping the room responsive. The stream is read on its own task, so audio and room events never wait behind a board. The queue holds two boards and the newest is the one that matters, because each board is the whole drawing rather than a change to it.
  • The one-second floor. The browser already debounces, but the agent still sends at most one board per second. A board that arrives too soon is kept as the latest, so read_board answers with it.
  • read_board. Returns JSON with the stroke count, the board's age and the timer, then resends the latest image, because a tool response cannot carry a picture.
  • Pause and silence. A paused interview drops boards the way it drops audio. Drawing counts as activity, so the silence nudge does not interrupt a candidate who is drawing without talking.

What the interviewer is told

There is one live prompt. The sentences that name the surface sit in a Surface struct with a coding and a whiteboard value, so a rule cannot end up in one prompt and not the other. Rebased onto #203, #192 and #202: the coding prompt is byte-identical to main's, and only the whiteboard entries are new in tests/golden/prompts.json. The thinking-time rule from #192 says "board changes" rather than "editor changes" at a board.

Phase checkpoints and grading

The report reviewer sees the board as it stood when each step closed, not just the final one. A candidate who clears the board between steps does not lose that evidence.

sequenceDiagram
    participant B as Browser
    participant A as Interview agent
    participant G as Gemini report model

    A->>B: framework_state (phases with evidence)
    B->>B: New phase ids, in REACTO order
    B->>A: JPEG + checkpoint id (after pointerup if mid-stroke)
    A->>A: Validate the id, keep one image per phase
    B->>A: Final JPEG when End is pressed
    B->>A: end_interview, once the uploads finish
    A->>G: generateContent(labelled inline JPEGs + prompt)
    Note over A,G: Retries and schema repairs resend every image
    G-->>A: Structured JSON report
    A-->>B: Report, stamped with interviewMode
Loading
  • Capture. Each framework_state update is reduced to known phase ids that have not been captured yet. A phase that completes while the pointer is down waits for pointerup, so the image and the replay describe the same stroke.
  • Labels. The agent validates the id again and assigns the label itself (Repeat, Example, Approach, Trace, Edge cases, Complexity). A string from the browser therefore never sits next to an image as an instruction to the reviewer.
  • Images sent to the reviewer. At most one image per phase, newest wins. The final board is added after the checkpoints when it differs from the last one.
  • Ordering. Uploads run one at a time, and end_interview is sent only after they finish. That way the report is never frozen on an older board.
  • Report prompt. It tells the reviewer what each phase means at a board, and says so when no board arrived.

Replay and recording

Replay stores no JPEGs. A board as a JPEG is over the per-event size limit and would use up the per-interview replay budget within a few frames. Replay instead stores the operations that drew the board: stroke, undo, redo and clear, in batches under 24 KiB, each optionally tagged with a checkpoint id. The replay page and the recording template rebuild any moment of the board with the same module the candidate drew on (createBoard + applyOp). Every operation is validated on the way back in: colour, width, point count and coordinates.

The report card

A whiteboard report opens on the board, below the committee summary:

  • "Your board, step by step." Each of the six steps shows the board as that step closed, the step's score, and what the interviewer recorded about it. A step with nothing recorded says so, rather than disappearing.
  • Final board. Shown after the six steps.
  • Names. Steps, Practice next and Framework evidence use the whiteboard names. The coding score is labelled Board work.
  • Schema. Unchanged. The card maps the REACTO ids it already carries to display names.
  • Images are not saved. The step images are the page's own copies of the checkpoints it sent. Six data URLs would be most of what a saved account report may weigh, so a report reopened from history says the pictures are in the recording.
  • Export. The .md file lists the steps and their evidence without base64.

Whiteboard

Whiteboard interview

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Need some time for testing and self-review; will mark this PR as ready for review once it's set.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Perhaps we will need to record the state of the whiteboard at each stage so that we can provide more information to the user at the end of the interview.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Additionally, pseudocode should be graded separately rather than being combined with optimization like it is currently.

@alanhc

alanhc commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

On the concern from #65 about whether Gemini can read a drawn board reliably: one option worth considering is Excalidraw as the board layer.

The drawing tools are not the point; the data model is. An Excalidraw scene is a list of typed elements, and boxes, arrows (bound to the boxes they connect) and text keep their actual strings. So next to the JPEG, the agent can get a plain-text summary of the board, e.g. box "left" -> box "mid", and typed pseudo-code arrives as exact text rather than something the model has to read off pixels. The report prompt can quote it too.

I checked whether it fits the vendoring rules, following the three-vrm.js precedent (one esbuild bundle, committed, with a reproduce recipe):

  • @excalidraw/excalidraw@0.18.1 + React 19 bundle into a single ES module: 4.9 MB (1.6 MB gzip), plus 145 KB of CSS. That is with @excalidraw/mermaid-to-excalidraw aliased to a stub; without it, 8.5 MB.
  • Fonts load from window.EXCALIDRAW_ASSET_PATH. The Latin ones are ~0.5 MB. The CJK font (Xiaolai) is 13 MB; I left it out, and Excalidraw then falls back to esm.sh for it. The current CSP refuses those loads (230 console violations in my run, drawing and export unaffected), but it needs either vendoring or silencing.
  • Served locally with every non-local request aborted, under the policy from src/web/policy.rs (script-src 'self', style-src 'self', nothing inline): it mounts in ~200 ms, freehand, shapes and text all work, and exportToBlob gives a ~20 KB JPEG. The only inline bit was setting EXCALIDRAW_ASSET_PATH, which moves to a same-origin script.
  • A 41-point freehand stroke is ~1.5 KB of JSON, so replay would still journal per-element changes from onChange rather than whole scenes.

The costs:

  • Far bigger than the 280-line whiteboard.js, and it brings React into the page.
  • The structured benefit exists only when the candidate uses shapes and the text tool. Freehand pseudo-code is still just points. Whether typed text belongs in a whiteboard interview is a product call: less realistic, much easier to grade.
  • Text input opens the integrity side: paste and library/file import would have to be disabled.

Only the board layer changes. The byte stream, read_board, the report attachment and the replay plumbing stay as they are. Happy to put together the vendored bundle and the scene-to-text summary, either on top of this branch or as a follow-up once it lands, whichever you prefer.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Hi, @alanhc

Thanks for the discussion!

Excalidraw was also brought up in the discussions under the Facebook post back then. After looking into it, I think it seems like a solid option to consider, but I also have a few thoughts.

If our goal is to simulate an actual interview environment as closely as possible, would Excalidraw make things too convenient for the user? After all, in a real interview, you only have a marker and a physical whiteboard. That raw experience is precisely what we want to deliver in this mode—allowing users to practice explaining their problem-solving approach and thought process purely through drawing on a blank board.

Currently, a tester, @Eason0729 , has tested this feature (the non-Excalidraw version) and provided a lot of feedback. Here is a brief summary of the key concerns raised from the testing:

  1. Should we support tablet touch input? Drawing with a mouse can significantly degrade the user's drawing experience and heavily impact their performance. Therefore, I would strongly prefer having this supported.

  2. Unreliable AI recognition will severely impact the quality of the questions asked.

For now, I'll focus on addressing the points raised in the feedback first. As for Excalidraw, we can discuss it further as a potential future improvement.

Overall, I actually think this is quite a promising direction. However, my primary concern is that users should have an experience that feels as close to reality as possible; we probably shouldn't compromise on that just to make recognition easier for the model.

WDYH

@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch 5 times, most recently from 2228cf7 to a124a50 Compare October 2, 2026 04:55
@ColtenOuO
ColtenOuO marked this pull request as ready for review October 2, 2026 05:16
cubic-dev-ai[bot]

This comment was marked as resolved.

@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from a124a50 to 2f6495a Compare October 2, 2026 22:15
cubic-dev-ai[bot]

This comment was marked as resolved.

@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from 2f6495a to 969a4bb Compare October 2, 2026 22:24
@alanhc

alanhc commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Agreed, realism should win for this mode. The point is practicing with a marker and a blank board, and shapes plus a text tool would take exactly that away, so I would drop Excalidraw as the board here. If it comes back at all, it would be as some separate "diagram" mode, not as the whiteboard.

On the two points from testing:

  1. Touch and pen. The board's pointerdown handler in web/interview.js returns early for anything but pointerType === "mouse", so touch and pen input is currently thrown away. Pointer events already unify the three, and touch-action: none is already set on the board, so most of the work may be lifting that check. The part that needs care is palm rejection: ignore touch pointers while a pen pointer is down, so a resting hand does not draw.

  2. Recognition. This does not need Excalidraw to soften. The candidate narrates while drawing, and that speech is already in the transcript, so the interviewer could be told to lean on what was said when the image is unclear, and to ask what a box or arrow stands for instead of guessing. That keeps the board raw and moves the ambiguity to a question a real interviewer would ask anyway.

Happy to help test the touch path once it is in.

jserv

This comment was marked as resolved.

@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch 3 times, most recently from c6fa1f4 to 9f0559b Compare October 4, 2026 19:54
@jserv
jserv requested a review from alanhc October 5, 2026 05:28
Comment thread src/agent.rs Outdated
Comment thread src/livekit/board.rs
Comment thread src/livekit/media.rs
Comment thread src/livekit/session.rs Outdated
Comment thread src/gemini.rs Outdated
Comment thread web/interview.js
Comment thread web/interview.js Outdated
Comment thread web/interview.js
Comment thread web/interview.js Outdated
Comment thread web/interview.js Outdated
Comment thread web/interview.js
Comment thread web/interview.js Outdated
@jserv

jserv commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Clearing the board between steps makes record_framework_evidence refuse every Coding, Test and Optimizations record for the rest of the interview, because the gate reads the newest board's stroke count. The report is then graded without that evidence. Details are on written_work in src/agent.rs.

@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from 9f0559b to a0d6bdb Compare October 5, 2026 07:08
@ColtenOuO
ColtenOuO requested a review from jserv October 5, 2026 07:40
Comment thread README.md
@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from a0d6bdb to 1947b11 Compare October 5, 2026 07:54
@jserv

jserv commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Measure the token cost of the whiteboard with the tooling from #203. Run a whiteboard interview and a regular one on the same problem, then compare them with scripts/analyze-gemini-usage.py. Report the first-generation promptTokenCount for each, and count the new instructions, tool declarations and board frames per text with countTokens, as #203 did. That tells us whether the cost is a fixed overhead or grows every turn.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Measured on six real interviews against gemini-3.1-flash-live-preview: problem two-sum, 20-minute coding + behavioral loop, editor and whiteboard alternating, three each, about five minutes each. Headless Chromium drove the page with a silent microphone; the editor arm typed code and the whiteboard arm drew eight figures, 30 seconds apart.

Opening Prompt tokens, first generation Runs
Editor 5,431 3 of 3 identical
Whiteboard 5,408 3 of 3 identical

Per text, with countTokens on gemini-3.1-flash-lite; the parts sum to the live difference (47 − 56 − 14 = −23):

Text Editor Whiteboard Difference Billed
System instructions 4,408 4,455 +47 every generation
Tool declarations 429 373 −56 every generation
Greeting 170 156 −14 every generation

Board frames are the part that grows. Each one is 60 image tokens on the Live socket regardless of content, and it stays in the context: by the end of each whiteboard session 8 to 10 boards were held, 480 to 600 tokens on every later generation until the sliding window trims them. Across a session that was 2.9 to 4.2% of the prompt tokens. A board goes out at most once a second after a one-second settle, or every four seconds while drawing continuously. countTokens prices the same JPEG as an inline image at 1,093.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

The results look fine in terms of cost.

@ColtenOuO
ColtenOuO requested a review from jserv October 5, 2026 09:23
@jserv

jserv commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

by the end of each whiteboard session 8 to 10 boards were held, 480 to 600 tokens on every later generation until the sliding window trims them. Across a session that was 2.9 to 4.2% of the prompt tokens. A board goes out at most once a second after a one-second settle, or every four seconds while drawing continuously. countTokens prices the same JPEG as an inline image at 1,093.

Good. Append the above report in docs/ directory.

A whiteboard interview takes the same problem bank and the same six
steps and swaps the editor and the test runner for a board. Nothing
runs, so correctness is what the candidate can defend by tracing an
example across their own drawing.

The board travels as a LiveKit byte stream rather than on a data topic,
because one board is tens of kilobytes and a data packet carries
fifteen, and it reaches Gemini the way a camera frame does, as a
realtime image. A tool response is JSON and cannot carry a picture, so
read_board asks for the board to be sent again and answers with what an
image cannot say: how much is on it and how long ago it was drawn.
A camera frame on that stream would displace the board, so a whiteboard
interview forwards no camera even where the operator enabled it.

Bundle 29, live prompt 21, follows the surface. A whiteboard
session is told it has no editor and no test runner, is offered
read_board in place of read_editor, and records board_snapshot where
the other records an editor snapshot or a test event; neither may
record the other's source. The phases about written work are gated on
strokes instead of on characters, and Test on the cases named against
the drawing instead of on a run, or an interview with no editor would
be refused the second half of its own flow. The report still reads the
transcript alone and the replay still carries no board.
The whiteboard was live but nowhere else: the report was written from a
transcript and an editor nobody opened, the recording kept no drawing,
and the checklist beside the timer called step four Coding while the
interviewer was asking the candidate to trace.

The board now rides the report request as an inline image, ahead of the
brief and on every repair, and the brief says what it is looking at:
nothing ran, so correctness is the trace the candidate walked, and the
phases keep their names with Coding meaning that trace, Test the cases
named against the drawing, and Optimizations the complexity they
confirmed. The system instruction is left as it is, so every report
call still shares its prefix, and the brief tells a whiteboard reviewer
how to read the rules that speak of code. That is report prompt 17,
inside the same bundle 29, and the reviewer is told when no board
arrived rather than sent looking for an attachment that is not there.

The recording keeps the drawing as the operations that made it. One
board as a JPEG is past the per-event ceiling on its own and would
spend the whole per-interview budget on a handful of frames; the same
board as strokes is a few kilobytes, so the replay page and the
recording template redraw any moment of the interview with the module
the candidate drew on, rather than the few moments a photograph could
afford.
A final board cannot represent work cleared between phases. Keep one
JPEG checkpoint per phase for grading and a vector marker for replay,
so review survives a clear without putting images into the replay
budget. Clear itself becomes one undoable step: with the earlier phases
kept, an accidental clear is the costliest mistake left at the board.

The checkpoints are why clearing is safe, and the gate on the phases
about written work has to agree: it read the newest board's stroke
count, so a clear refused every later Coding, Test and Optimizations
record. It now asks whether any board carried a drawing, and the count
leaves out erasures, which leave no ink. A board stream that never
finishes gives its read slot back after a deadline, and the image
read_board asks for reaches Gemini ahead of the tool response, which
Gemini starts answering as soon as it arrives.
One fixed width was either too wide to take out a character or too
narrow to clear a region. Three offered sizes cover both, and a fixed
list rather than a slider keeps every width on the replay one a button
can make. The cursor becomes a circle the size of the erasure.
A whiteboard report read as an editor one: a Coding score, REACTO step
names the candidate was never shown, and one board folded at the foot
of the card. The page already captured the board as each step closed,
for the grader, and kept none of it.

The report now opens on the board, step by step: the image as that step
closed, its score and what the interviewer wrote down, then the final
board. Steps, plans and evidence take the whiteboard names, and the
score is Board work. Only the display changes; the report keeps the
REACTO ids it is graded against. The images are never saved, so a
report reopened from history says where they are instead.

Those step images only exist if every board reaches the agent before
the report is frozen. The board stays locked until the room is joined,
so the replay and the interviewer see its first stroke. A reconnect
sends boards ahead of a queued end_interview, an end mid-stroke closes
the stroke and captures the phase parked behind it, and the end waits
for uploads only briefly. A checkpoint that failed to upload is retried
ahead of every later board, not only on a reconnect, and a failed write
still closes its stream. Steady drawing sends a board every few seconds
instead of waiting for a pause that never comes.
The board threw away every pointer that was not a mouse, so a tablet
could not draw on it at all. Pointer events already carry all three, and
the canvas already sets touch-action: none, so the board now takes any
of them, one pointer at a time. A hand resting on a tablet reports as
touch: a touch is refused while a pen is down, and a pen that lands
after a palm cancels the palm's stroke rather than leaving it drawn.
A stroke that begins also cancels the settle timer the previous one
started, so a board is never exported half-drawn.
A board reaches the interviewer as an image of handwriting and sketches,
and a box or arrow it cannot read is one it fills in with a guess. The
candidate narrates while drawing, and that speech is already in the
conversation, so a whiteboard session is told to read an unclear mark by
what was said while it was drawn, and failing that to ask what it stands
for, as a person at the board would. The board stays raw and the live
prompt keeps version 21, which has not shipped.
The README said what the whiteboard sends and where, but not how a
candidate picks it, draws on it or reads its report, so a first try had
to discover the toolbar, the input rules and the renamed steps alone.
A section now walks through choosing it in the lobby, when the board
takes strokes, pen and touch input, each toolbar control, what Jim sees
and when, the six board steps, and what the report shows.
Measured on six real interviews, three per surface, so a reviewer can
tell the fixed overhead from the part that grows: the opening is 23
tokens below the editor's, and each board sent adds 60 image tokens
that every later generation is billed on again.
@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from 1947b11 to 03d0222 Compare October 5, 2026 09:30
@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Good. Append the above report in docs/ directory.

added in 03d0222

@jserv
jserv merged commit ef8823c into sysprog21:main Oct 5, 2026
13 checks passed
@jserv

jserv commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Thank @ColtenOuO for contributing!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a whiteboard interview mode

3 participants